Argument co-occurrence matrix as a description of verb valence

نویسندگان

  • Łukasz Dębowski
  • Marcin Woliński
چکیده

A new description of verb valence is proposed. Rather than by full valence frames, verb subcategorization is described by the list of arguments and the matrix of five distinct pairwise argument interactions. This approach is argued to be computationally robust and sufficient although it has been initially motivated only by the need of a reliable test set for machine verb valence learning. Speaking more abstractly, we propose an approach to decomposing empirical subsets of powersets that reduces data sparseness problems and an approach to discretize (summarize) a large number of contingency tables. Both techniques are novel and might be useful not only in linguistics.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Study of Valence & Argument Integration in Chinese Verb-Resultative Complement

Verb resultative complement (VC) is a common structure of Chinese language with abundant forms of collocation. It makes much sense for VC research to analyze the general rules of argument integration in light of diversities of predicate & complement and the complexity of argument integration in the forming of VC with predicate & complement. This article has analyzed and summarized the existing ...

متن کامل

The Valence Patterns of Japanese Verbs Extracted From The EDR Corpus

This paper describes research on particular verb valences obtained from actual linguistic data. We created verb valence data using data from the EDR Co-occurrence Dictionary as our source. The EDR Co-occurrence Dictionary is coded with syntactic governing-dependent relation tags and semantic tags. The syntactic governing-dependent relations data in the EDR Co-occurrence Dictionary however, is e...

متن کامل

Valence extraction using EM selection and co-occurrence matrices

This paper discusses two new procedures for extracting verb valences from raw texts, with an application to the Polish language. The first novel technique, the EM selection algorithm, performs unsupervised disambiguation of valence frame forests, obtained by applying a non-probabilistic deep grammar parser and some post-processing to the text. The second new idea concerns filtering of incorrect...

متن کامل

Modeling Subcategorization through Co-occurrence: a Computational Lexical Resource for Italian Verbs

1. Goals and Methodology The aim of this abstract is to introduce LexIt, a freely available lexical resource to characterize Italian verb argument properties in terms of distributional information automatically extracted from large corpora with state-of-the-art computational linguistics methods. Research on automatic extraction of subcategorization frames from corpora has a long tradition in co...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2007